Data quality in OSH

Data quality in occupational safety and health (OSH) refers to the suitability of records to the preventive decisions they are meant to support, including their accuracy, coverage, consistency, timeliness, and traceability.

In short

Useful data should be representative of the work and interpretable within its context. Simply filling in fields is insufficient if the definitions, population, or sources are not comparable.

Content
  1. Quality linked to a purpose
  2. Definitions and master data
  3. Accuracy, coverage, and timeliness
  4. Consistency of calculations and indicators
  5. Controls during the life cycle
  6. Personal data and artificial intelligence
  7. Practical example
  8. Quality monitoring
  9. Related concepts
  10. On the blog
  11. References

AZ Dictionary →

Quality linked to a purpose

The quality of a dataset depends on its intended use. A record sufficient to locate a document may be insufficient to compare exposures between centers. Before collecting information, it is advisable to define what decision is being supported, what population is being analyzed, and what level of detail is required.

In occupational health and safety (OSH) analytics, problems often arise when combining records created for different purposes. Accidents, inspections, training, and health surveillance do not automatically share the same units or time periods. Integration must preserve the meaning of each data point so that a numerical result does not appear more robust than it actually is.

Definitions and master data

The organization needs common criteria for centers, positions, tasks, teams, people, and types of events. It’s advisable to document what each field means, who maintains it, and when it changes. The same word can be used differently between units; for example, a measure might be marked as closed when equipment has been purchased or only when its effectiveness has been verified.

Stable identifiers allow records to be linked without confusing similar-sounding names. Historical context must also be preserved: if a person changes positions, the previous accident remains linked to the conditions that existed when it occurred. Retrospectively replacing all data with the current situation can distort the analysis.

Accuracy, coverage, and timeliness

Accuracy requires that the information correctly represents the relevant fact or attribute. Coverage allows us to know which cases are included and which are missing. Timeliness refers to the data being available in time for decision-making. These dimensions can fail independently: a correct record may arrive too late or represent only part of the work.

The absence of a recorded accident does not necessarily equate to the absence of accidents. Similarly, a facility with few observations may have fewer problems or report them less. The review should cross-reference sources, facilitate reporting, and analyze gaps before presenting a comparison as evidence of improved performance.

Consistency of calculations and indicators

OSH indicators require a compatible numerator, denominator, period, and scope. If events involving both company staff and contractors are included, a population representing only company staff should not be used without explanation. The hours, people exposed, or tasks observed must correspond to the phenomenon being compared.

Leading and lagging indicators offer different perspectives. The number of inspections reflects activity, while the verified correction of findings provides further evidence. It is advisable to avoid undocumented changes in definition, rounding that masks differences, and comparisons of small groups without explaining their variability.

Controls during the life cycle

Controls can include required fields, valid formats, duplicate detection, and consistency rules. They should be designed to facilitate accurate recording without forcing fictitious responses when the information is not yet available. It is preferable to identify missing data and assign it for review rather than turning uncertainty into a seemingly definitive value.

Documented information requires assigned responsibilities, version control, and a sufficient change history. Corrections should be propagated to affected reports in a controlled manner. It’s also advisable to review imports and connections between systems: an entire table may contain errors if units, dates, or identifiers were exchanged during the transfer.

Personal data and artificial intelligence

Quality must be coordinated with purpose, minimization, and data protection. Collecting all available information is not an appropriate criterion. Health data require specific safeguards and restricted access; general analyses should use the level of aggregation necessary for their purpose and avoid unnecessarily identifying individuals.

In the use of AI, in addition to individual accuracy, the suitability of the whole is important. The data must represent the situations where the system will be used. A model may appear to work well with historical records but fail when faced with a new process, a different population, or a change in the reporting method.

Practical example

In a hypothetical example, two centers compare the percentage of training completed. One includes all active staff, while the other also includes people who no longer work there. Furthermore, the first requires a final evaluation, while the second considers simply offering the course sufficient. The difference in the indicator reflects issues of population and definition.

The organization agrees on the criteria, corrects the scope, and documents the date of the change. It recalculates comparisons that can be redone with reliable information and identifies limitations of the remaining historical data. The result is not just cleaning up a table: it establishes rules to prevent the same problem from recurring in the next report.

Quality monitoring

It’s important to measure relevant problems: duplicate records, missing essential fields, delays, discrepancies, and recurring corrections. Each quality indicator needs someone responsible who can address the root cause of the error. Manually correcting a report every month can mask a persistent problem with data capture or integration.

The protection of health data and its preventive value must be maintained throughout the entire process. The purpose of quality is to enable better-informed decisions by clarifying what is known and what is not. A reliable system preserves the limitations of the data rather than masking them with more charts or figures.

Related concepts

On the blog

References

  1. International Labour Organization. How data can protect workers’ health and lives. 2017. Official source
  2. International Labour Organization. Resolution on statistics of occupational injuries caused by work-related accidents. 1998. Official source
  3. National Institute for Occupational Safety and Health. NTP 1211: Accident statistics in the company. 2024. Official source
  4. Spanish Data Protection Agency. Accuracy, suitability and quality of data in personal data processing with AI. 2026. Official source
  5. European Union. Regulation (EU) 2016/679, General Data Protection Regulation. Official source

Editorial information

Publication date: October 10, 2026.

Editorial Manager: Sabentis Editorial Team.

Author: Pablo Rodríguez LinkedIn

Executive Vice President of the ORP International Foundation and Chief Financial Officer of Sabentis.

Request a Demo

Discover all that Sabentis can do for your organization.

Try Sabentis

request a demo
stars 5
GetApp Software Advice Capterra